Goto

Collaborating Authors

 advanced task


Game-Time: Evaluating Temporal Dynamics in Spoken Language Models

arXiv.org Artificial Intelligence

Conversational Spoken Language Models (SLMs) are emerging as a promising paradigm for real-time speech interaction. However, their capacity of temporal dynamics, including the ability to manage timing, tempo and simultaneous speaking, remains a critical and unevaluated challenge for conversational fluency. To address this gap, we introduce the Game-Time Benchmark, a framework to systematically assess these temporal capabilities. Inspired by how humans learn a language through language activities, Game-Time consists of basic instruction-following tasks and advanced tasks with temporal constraints, such as tempo adherence and synchronized responses. Our evaluation of diverse SLM architectures reveals a clear performance disparity: while state-of-the-art models handle basic tasks well, many contemporary systems still struggle with fundamental instruction-following. More critically, nearly all models degrade substantially under temporal constraints, exposing persistent weaknesses in time awareness and full-duplex interaction. The Game-Time Benchmark provides a foundation for guiding future research toward more temporally-aware conversational AI. Demos and datasets are available on our project website https://ga642381.github.io/Game-Time.


Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

arXiv.org Artificial Intelligence

Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferret-UI 2, a multimodal large language model (MLLM) designed for universal UI understanding across a wide range of platforms, including iPhone, Android, iPad, Webpage, and AppleTV. Building on the foundation of Ferret-UI, Ferret-UI 2 introduces three key innovations: support for multiple platform types, high-resolution perception through adaptive scaling, and advanced task training data generation powered by GPT-4o with set-of-mark visual prompting. These advancements enable Ferret-UI 2 to perform complex, user-centered interactions, making it highly versatile and adaptable for the expanding diversity of platform ecosystems. Extensive empirical experiments on referring, grounding, user-centric advanced tasks (comprising 9 subtasks $\times$ 5 platforms), GUIDE next-action prediction dataset, and GUI-World multi-platform benchmark demonstrate that Ferret-UI 2 significantly outperforms Ferret-UI, and also shows strong cross-platform transfer capabilities.


Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

arXiv.org Artificial Intelligence

Recent advancements in multimodal large language models (MLLMs) have been noteworthy, yet, these general-domain MLLMs often fall short in their ability to comprehend and interact effectively with user interface (UI) screens. In this paper, we present Ferret-UI, a new MLLM tailored for enhanced understanding of mobile UI screens, equipped with referring, grounding, and reasoning capabilities. Given that UI screens typically exhibit a more elongated aspect ratio and contain smaller objects of interest (e.g., icons, texts) than natural images, we incorporate "any resolution" on top of Ferret to magnify details and leverage enhanced visual features. Specifically, each screen is divided into 2 sub-images based on the original aspect ratio (i.e., horizontal division for portrait screens and vertical division for landscape screens). Both sub-images are encoded separately before being sent to LLMs. We meticulously gather training samples from an extensive range of elementary UI tasks, such as icon recognition, find text, and widget listing. These samples are formatted for instruction-following with region annotations to facilitate precise referring and grounding. To augment the model's reasoning ability, we further compile a dataset for advanced tasks, including detailed description, perception/interaction conversations, and function inference. After training on the curated datasets, Ferret-UI exhibits outstanding comprehension of UI screens and the capability to execute open-ended instructions. For model evaluation, we establish a comprehensive benchmark encompassing all the aforementioned tasks. Ferret-UI excels not only beyond most open-source UI MLLMs, but also surpasses GPT-4V on all the elementary UI tasks.


The Simple ML release and its big data implications for Sheets users

#artificialintelligence

Last week, Google announced and released a beta version of Simple ML for Sheets, a TensorFlow Decision Forests-produced add-on for Google Sheets. This release is one of the first of its kind, offering many simple and some complex machine learning functionalities directly to Google Sheets users. Although Simple ML has been touted as the machine learning solution for people with no prior knowledge of machine learning, the Advanced Tasks it offers promise value to data scientists, machine learning experts and anyone else working with bigger datasets. Read on to learn more about this release and how it may shape spreadsheet-based data and machine learning projects in the future. Simple ML for Sheets is currently available in beta.


Robots will build better jobs

#artificialintelligence

Computers, artificial intelligence programs, and robots are doing more of the work that humans used to do. Data from the Bureau of Economic Analysis show that the amount of technology used per unit of production doubled between 2001 and 2015.* There's no reason to believe this trend will slow anytime soon. What does that mean for jobs? Researchers from the University of Oxford estimate that 47% of U.S. jobs could be automated as soon as 2025, meaning about 70 million Americans could find themselves out of work.**